Why maths at all?
Before we go any further, let me address the question that every beginner has at this point: do I actually need to understand this?
The honest answer is yes, but not in the way you might fear. You do not need to perform matrix multiplication by hand on an exam. What you do need is a mental model for what is happening when your code runs. When you understand that an AI model is essentially doing an enormous number of vector operations, the behaviour of the model starts to make sense in a way it simply cannot if you treat it as a black box.
Here is the key insight: vectors and matrices are just ways of organising numbers. That is it. All the intimidating notation is just a shorthand for things you already understand.
"The miracle of the appropriateness of the language of mathematics for the formulation of the laws of physics is a wonderful gift which we neither understand nor deserve."
Eugene Wigner, physicistA vector is just a list of numbers
Forget arrows. Forget geometry for now. At its most practical level, a vector is simply an ordered list of numbers. That is it. If you have ever looked at a row in a spreadsheet, you have worked with a vector.
This vector represents one Titanic passenger: ticket class 3, aged 22, fare £7.25, 0 siblings aboard, and 1 parent aboard.
Every row in your dataset becomes a vector. Every observation is a point in a multi-dimensional space, with one dimension per feature.
The number of elements in a vector is its dimension. The passenger vector above has 5 dimensions. A word embedding vector in a language model like GPT might have 768 or even 4,096 dimensions. You cannot visualise that space, but the maths works exactly the same way.
Your location on Earth can be described by exactly two numbers: latitude and longitude. That is a 2-dimensional vector. If you add altitude, it becomes 3-dimensional. An AI model describing the meaning of a word might use 768 numbers, forming a 768-dimensional vector. More dimensions means more nuance, more information, more expressiveness.
A matrix is a table of numbers
A matrix is simply a grid of numbers arranged in rows and columns. When you look at your whole dataset, all your passengers and all their features together, you are looking at a matrix. Rows are observations. Columns are features.
This is a 5×5 matrix. Shape: 5 rows by 5 columns. The highlighted row is one vector representing one passenger. When you call df.values in Pandas, you get exactly this.
Every matrix has a shape: (rows, columns). In NumPy you will write X.shape constantly.
Images are matrices
You saw in Lesson 2.1 that a grayscale image is a grid of numbers. That grid IS a matrix. A 28×28 pixel grayscale image is a 28×28 matrix, which is 784 numbers in total. A colour image adds a third dimension: 28×28×3 (one channel for red, green and blue). This 3-dimensional structure is called a tensor.
This is why the main AI framework is called TensorFlow and the operations are called tensor operations. Everything is just numbers in arrays of various shapes. The word "tensor" sounds intimidating but it is just a generalised word for a multi-dimensional array of numbers.
In practical AI work, most bugs and errors come from mismatched shapes. You try to multiply two matrices but their dimensions do not align. You feed data with the wrong shape into a model. Understanding that every piece of data has a shape, and that shapes must be compatible for operations to work, will save you enormous frustration when you start coding.
Where vectors and matrices are used in AI
The dot product: AI's most important operation
There is one vector operation that appears more than any other in AI: the dot product. It takes two vectors of the same length and returns a single number. Here is why it matters.
= (2×4) + (3×1) + (1×5)
= 8 + 3 + 5
This simple operation powers an enormous amount of AI. In neural networks, every layer computes a dot product between your data vector and a weight vector. In transformers, attention is computed using dot products between query and key vectors. In recommendation systems, similarity between a user vector and a product vector is computed with a dot product.
You do not need to compute this by hand. NumPy does it in one line: np.dot(a, b). But understanding what it means, asking "how similar are these two things?", is the insight that will help you understand why AI works the way it does.